Papers by Barrett Martin Lattimer
Sparse Rewards Can Self-Train Dialogue Agents (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models have been driven by supervised fine-tuning and high-quality human feedback. however, acquiring meaningful human feedback has become increasingly challenging and costly. |
| Approach: | They propose a method that empowers LLM agents to enhance their performance without external feedback. |
| Outcome: | The proposed method improves tool-based interactions while preserving general model capabilities across diverse benchmarks. |